Journal of Medical Internet Research
◐ JMIR Publications Inc.
Preprints posted in the last 7 days, ranked by how well they match Journal of Medical Internet Research's content profile, based on 87 papers previously published here. The average preprint has a 0.11% match score for this journal, so anything above that is already an above-average fit.
Pavia, M. J.; Amaro, I. F.; Xu, D.; Gonzalez-Hernandez, G.; Scotch, M.
Show abstract
Influenza vaccine effectiveness (VE) is estimated from a limited number of clinics using a test-negative design. These standard estimates face geographic, temporal, and operational constraints. Using Twitter/X data, we applied few-shot chain-of-thought prompting to identify self-reported vaccination status and influenza test results, then implemented a test-negative-like design to estimate VE. Our estimates fell within the range of interim reports and could complement current systems, improving feasibility, timeliness, and scalability.
Sierpe, A.; Yen, R. W.; Milliman, A.; Cady, E.; Ahn, B.; Dade, A. E.; Devito, A. M.; Eckert, B. A.; Gopalan, V. V.; Krasinski, S. C.; MacMartin, M. A.; Musacchio, S. G.; Zhang, J.; Saunders, C. H.
Show abstract
Background Agenda-setting is a fundamental patient-centered communication practice in which a clinician works with a patient to elicit, propose, and organize topics for discussion during a clinical encounter. Various agenda-setting interventions have been developed, including patient-facing tools and clinician training, but their effects have not been systematically evaluated. We aimed to determine the effects of these interventions on encounter, patient, care partner, and clinician outcomes. Methods We searched grey literature and seven databases, including PubMed, from inception through July 2025 for randomized and non-randomized comparative studies of interventions designed to promote or improve clinical visit agenda-setting. Two reviewers independently screened articles and extracted data, with a third reviewer resolving conflicts. We assessed risk of bias using RoB 2 for randomized studies and ROBINS-I for non-randomized studies. We conducted random effects meta-analyses when outcomes were sufficiently comparable, assessed heterogeneity using I2, and rated certainty of evidence using GRADE. Post hoc exploratory subgroup analyses examined study design, adjustment status, and intervention structure. Results Twenty-nine articles describing 22 unique studies met the inclusion criteria, including 13 randomized and nine non-randomized studies. Agenda-setting interventions increased the occurrence of agenda-setting (risk ratio 5.43, 95% confidence interval (CI) 2.06 to 14.28, I2=34.6%) and favored the intervention for concerns addressed when measured as a continuous outcome (standardized mean difference (SMD) 0.37, 95% CI 0.16 to 0.57, I2=65.3%) and overall clinician satisfaction (SMD 0.50, 95% CI 0.23 to 0.78, I2=0.0%). There were no clear differences in the number of concerns raised (mean difference (MD) 0.21, 95% CI -0.19 to 0.61, I2=59.6%), visit duration (MD 0.64 minutes, 95% CI -0.83 to 2.12, I2=51.4%), or overall patient satisfaction (SMD 0.05, 95% CI -0.05 to 0.15, I2=47.0%). Potentially important heterogeneity was present for four of these six outcomes. Post hoc exploratory subgroup analyses did not provide clear evidence that effects varied by study design, adjustment status, or intervention structure. Risk of bias was often high, serious, or critical, and certainty of evidence was low or very low for all pooled outcomes. Conclusions To our knowledge, this is the first comprehensive synthesis of clinical visit agenda-setting interventions. Such interventions may increase the occurrence of agenda-setting and the extent to which patient concerns are addressed without increasing visit length. However, the certainty of evidence was low or very low, and the available evidence does not establish a superior intervention structure.
Edmond, E. C.; Dreyer, A. J.; Winston, A.; Khoo, S. H.; Joska, J.; Nightingale, S.
Show abstract
Background Computerised cognitive testing may address the global challenge in identifying cognitive changes in people living with HIV scalably and affordably. We assessed a computerised battery (CB) of cognitive tests, in a prospective cohort (CONNECT) of people with HIV in a low-income peri-urban area of Cape Town, South Africa during a national programmatic switch from efavirenz- to dolutegravir-based antiretroviral therapy (ART). Methods We recruited 170 people with HIV and 91 people without HIV (controls) (140[82%] and 41[45%] followed up). The CB and gold-standard pen&paper cognitive testing (P&P) were performed at both timepoints. Technology familiarity/use questionnaire data were also collected. We compared performance in detecting lower group-level cognitive performance associated with efavirenz treatment. Furthermore, the CB was compared to P&P in classifying individuals with low cognitive performance, correlation of global test scores and domain-level scores between batteries, and practice effects between timepoints. Exploratory principal component analysis was also performed. Results People with HIV on efavirenz at baseline had lower performance on the computerised battery than controls, {Delta}T=2.6, p=0.0047. This difference was lost after switching to dolutegravir-based ART at follow-up. CB and P&P global T were moderately correlated (R2=0.203, p<0.001), and the CB performed moderately in classification of low cognitive performance against the gold standard (AUC 0.70, sensitivity 0.52, specificity 0.76, PPV 0.40, and NPV 0.84). Selecting the first three principal components improved both classification of low cognitive performance (AUC 0.77) and correlation strength with P&P global T (R2=0.3, p<0.001). The CB did not show practice effects. Most participants owned a mobile phone (95%, 85.9% of these smartphones). Performance was better in smartphone owners ({Delta}T=1.8) and computer owners (23%, {Delta}T=1.8). Conclusions Delivering computerised cognitive testing was feasible in this low-income southern African setting. The CB showed reasonable construct validity (detecting known lower cognitive performance associated with efavirenz-ART) and may detect broad cognitive characteristics such as processing speed and accuracy. However, correlation of CB results with gold standard P&P testing was low-moderate and may limit its applicability as a diagnostic tool. This might be improved by including a wider range of cognitive domains tested in the CB, or data driven analysis. Brief CBs may fulfil an initial screening role, followed by more detailed clinical assessment.
Chin, A. T.; Zhu, N.; Vangala, S.; Woo, H.; Wisk, L. E.; Kingsley, T.; Mafi, J. N.; Lukac, P. J.
Show abstract
BACKGROUND Generative AI (genAI) chart summarization tools embedded in electronic health records (EHRs) are being rapidly deployed across U.S. health systems. Although these tools represent a promising solution to alleviate cognitive burdens, their effects have not been examined in randomized-clinical trials (RCTs). METHODS In this pragmatic RCT at a single academic health system, 284 outpatient clinicians across forty-two specialties were assigned 1:1 to Epic's outpatient chart summarization tool or a usual-care control arm over 90 days, from February 23 to May 23, 2026. The primary outcome was physician task load (PTL) adapted for pre-charting. Prespecified exploratory outcomes included additional validated psychometrics as well as usability, safety, and time-based measures. Descriptive statistics included interaction and usage of the tool. RESULTS Of 74,474 AI chart summaries generated, 14.2% were interacted with by a clinician; the proportion of generated summaries interacted with declined from 21.5% in month 1 to 10.5% in month 3, and the proportion of clinicians using the tool at least once per month declined from 88.7% to 66.2%. The adjusted between-arm difference in PTL at follow-up favored the intervention arm (scale 0-400; -27.4; 95% CI, -49.4 to -5.3; P=0.02). Among the Professional Fulfillment Index (PFI; scale 0-4, lower=better) psychometrics, overall burnout (-0.20; 95% CI, -0.38 to -0.01) and work exhaustion (-0.24; 95% CI, -0.47 to -0.02) were lower in the intervention arm, with little difference in overall professional fulfillment (+0.04; 95% CI, -0.16 to 0.25). Charting time per encounter showed no significant between-arm difference during steady state (-1.2 seconds; 95% CI, -19.0 to 16.6). The net promoter score was -22, indicating that on average, clinicians did not recommend the tool. Among free-text respondents, 57.1% reported at least one concern, most commonly tool limitations or inaccurate information. No adverse patient safety events or near-misses were reported. CONCLUSION An EHR-integrated AI chart summarization tool modestly reduced physician task load and was associated with lower burnout, without time savings and against declining engagement. Sustained usage and oversight of reported inaccuracies remain open challenges.
Wojcik, S.; Rulkiewicz, A.; Domienik-Karłowicz, J.
Show abstract
Large language models perform well on medical examinations, but users routinely challenge their answers and invoke professional roles, and it is unclear what a system does when a medical credential and a stated task-specific accuracy point in opposite directions. In a factorial experiment on 480 items from four Polish specialty examination sets and three consumer large language model systems (ChatGPT, Claude, Gemini), each item and system received eleven independent conversations. Conditions crossed attributed source role (medical student, experienced specialist), stated prior accuracy on similar questions (2/10, 8/10) and suggestion correctness. The primary outcome was adoption of a prespecified incorrect option when the baseline answer matched the official key, comparing a specialist described as 2/10 with a student described as 8/10. Baseline agreement with the key was 87.2% across 15,683 analyzable conversations. The incorrect option was adopted more often from the specialist described as 2/10 than from the student described as 8/10 (10.2% vs. 7.6%; adjusted risk difference +2.82 percentage points, 95% CI +0.65 to +4.99). Estimates varied across the three systems and only one system-specific interval excluded zero. In a prespecified exploratory analysis with a shared eligibility rule, correct suggestions were adopted far more often than incorrect ones (risk difference +35.7 percentage points, 95% CI +30.8 to +40.7), indicating selective rather than indiscriminate compliance. An incorrect suggestion from a specialist with low stated accuracy was therefore slightly more influential than the same suggestion from a student with high stated accuracy, although the difference was modest and varied across systems. Agreement reached only after a user has disclosed a preferred answer should not automatically be treated as an independent second opinion, and medical large language model systems should be evaluated on how they revise answers after such disclosure, not solely on initial accuracy.
Kohler, S.; Meyer-Eschenbach, F.; Michelena, X.; Marschollek, M.; Eils, R.
Show abstract
The openEHR standard provides an open, vendor-neutral architecture for clinical data repositories (CDRs), yet its real-world deployment has not been systematically documented. We conducted a dual-perspective survey combining a vendor survey of openEHR CDR providers with a community survey of openEHR practitioners. Eleven vendor organisations reported deployments across 22 countries and over 100 institutions and health regions. A complementary community survey (n=29, 17 countries) provided context on regulatory environments, adoption drivers, and barriers. Combined, the surveys cover 28 countries, 26 of them with a reported openEHR CDR deployment. Three findings emerge: openEHR has achieved national-scale presence through two distinct channels. Through vendor-market convergence, openEHR-based systems cover the majority of regional health authorities without a national mandate, including 19 of 21 Swedish regions, 3 of 4 Norwegian health regions, and 16 of 21 Finnish wellbeing services counties. Through national health record adoption, governments have built or procured national systems on openEHR as their technical foundation, including Ireland, Malta, Greece, Jamaica and Slovenia. Across Europe, this constitutes an openEHR-based interoperability infrastructure already in place across multiple EU member states. We identified no country in which openEHR is named in binding national regulation, creating structural fragility and an unrealised opportunity for alignment with the European Health Data Space (EHDS). Second, 61% of deployments serve primary use only, and 12% support both primary and secondary use. Third, lack of openEHR-specific knowledge is the most consistent adoption barrier across all geographies and deployment tiers. Adoption is driven by practitioner need and innovation, not by regulatory mandate.
Wain, K. F.; Carroll, N. M.; Maclennan, A. J.; Hixon, B.; Steiner, J.; Ritzwoller, D. P.
Show abstract
Purpose: Lung cancer screening (LCS) with low-dose computed tomography (LDCT) reduces lung cancer mortality, yet screening participation remains low. We evaluated whether a brief informational video nudge delivered immediately before a scheduled clinical encounter increased LCS ordering and baseline LCS completion. Patients and Methods: We conducted a randomized feasibility trial within Kaiser Permanente Colorado from March through October 2025. LCS-eligible patients with an upcoming primary care or pulmonology appointment were assigned to intervention or usual care based on birth month. Intervention patients were split into two group, a group who received the LCS informational video nudge via text message within 24 hours of an eligible appointment; and second group who received the text plus a QR code video link during appointment rooming. Outcomes included LCS orders, baseline LCS-LDCT completion, and video engagement. Multivariable logistic regression was used to evaluate factors associated with LCS ordering. Results: Among 1,093 patients, 549 were assigned to intervention and 544 to usual care. Intervention patients were more likely to receive an LCS order within 1 day of their appointment (22.6% vs 16.4%; p=.010) and any time during follow-up (32.6% vs 24.1%; p=.002). Baseline LCS-LDCT completion was 51% higher in the intervention group, although the difference was not statistically significant (8.6% vs 5.7%; p=.078). Among the intervention group, 93 individuals (17%) viewed the video, generating 114 total views, and viewers watched an average of 79% of the video. Most views (82.5%) occurred through text-message delivery rather than QR codes. Conclusion: A brief, low-burden LCS informational video delivered immediately before a clinical encounter and integrated into existing workflows significantly increased LCS ordering and was associated with higher screening completion. Timely, scalable digital nudges may provide an effective strategy for improving LCS participation. Based on the observed effectiveness, feasibility, and efficiency of the intervention, KPCO incorporated the behavioral nudge into standard clinical care in February 2026.
Hickman, R.; Joyce, D. W.; Gray, N.; Shergill, S.; D'Oliveira, T. C.
Show abstract
Background: Shiftwork disrupts natural sleep-wake cycles, alters light exposure patterns, and contributes to circadian misalignment. Detrimental health consequences associated with shift work include elevated risk for metabolic disorders, cardiovascular disease, cancer and all-cause mortality. Healthcare workers have one of the highest rates of shift work exposure, yet there are relatively few non-pharmacological interventions (with good evidence) developed to improve sleep outcomes in this population. Objective: A pre-post pilot interventional study assessed the acceptability and perceived effectiveness of commercial noise-masking earbuds on improving subjective sleep characteristics among National Health Service (NHS) healthcare staff working fast rotating shifts. Methods: Noise-masking sleep earbuds (Kokoon NightBuds) were worn for a pilot six-week intervention by twenty-seven NHS nurses (aged 26-43 years, 88.9% female) working fast rotating shifts from the EClocker Study. Sensors inside the earbuds were paired with a smartphone app to monitor sleep. An audio library in the smartphone app delivered personalised relaxation exercises and sleep techniques drawn from cognitive behavioural therapy for insomnia (CBT-I). A pre-post two-week monitoring period with daily smartphone-based Experience Sampling Methods (ESM) captured perceived daily sleep patterns. Acceptability and perceived effectiveness of the earbuds in promoting better sleep outcomes was assessed. Results: Use of the noise-masking sleep earbuds over a six-week period was associated with positive sleep improvement trends and elicited promising acceptability. Almost two thirds of NHS fast rotating shift nurses (63%) subjectively reported reductions in general sleep disturbance symptoms (PSQI Global), one in four experienced perceived sleep quality improvements (SQ; 25.9%), one in five reported sleeping longer (TST; 22.2%), and a third perceived falling asleep faster (SOL; 33.3%), had better sleep efficiency (SE; 33.3%) and improved daytime dysfunction (33.3%) (PSQI subcomponent scores). Sleep diaries (CSD) collected daily using smartphone-based ESM also demonstrated small improvements post-sleep earbud use; nurses reported sleeping an average 18 minutes longer (TST) and fell asleep more easily, on average 11 minutes faster (SOL). Sleep earbuds were generally well tolerated; 56% of nurses reported the earbuds as (somewhat to very) helpful, 52% reported (somewhat to strongly) falling asleep more easily (SOL), 44% felt (somewhat to strongly) their sleep quality was improved (SQ) and 30% agreed (somewhat to strongly) they slept longer (TST) and had less disturbed sleep. Conclusions: To our knowledge, this is the first study in Europe to pilot noise-masking earbuds as a potential non-pharmacological aid to improve sleep-wake behaviours or mitigate fatigue for healthcare staff. Preliminary results showed promising acceptability and (small) perceived sleep improvement trends following a targeted six-week earbud intervention in NHS fast rotating shift nurses.
Chowdhury, A. R.; Chowdhury, B.
Show abstract
Background: Consumer use of AI chatbots for health advice is rising, yet triage safety relative to established services remains unclear. Australia's Healthdirect, a government-backed symptom checker with 2.4 million uses in FY2024-25, remains unevaluated against frontier large language models (LLMs), and whether premium subscriptions improve triage safety remains unexplored. This study compared the triage accuracy and safety of Healthdirect against six LLM configurations across ChatGPT, Claude, and Gemini, assessed whether paid subscriptions improve triage safety, and characterised each system's error patterns. Methods: Forty-five clinical vignettes from the Semigran et al. benchmark spanning emergency, non-emergent, and self-care categories (15 each) were evaluated across seven systems. Healthdirect was tested following a seven-rule interaction protocol. LLMs were evaluated using first-person patient-language prompts under free-tier and paid-tier conditions. Outcomes were triage accuracy, emergency sensitivity, under-triage, and critical misses, analysed using Cochran's Q, Bonferroni-corrected McNemar tests, Cohen's kappa, and Wilson intervals. Findings: Triage accuracy differed significantly (Cochran's Q = 36.79, p < 0.001). Healthdirect achieved 48.9% accuracy (95% CI 35.0% to 63.0%; kappa = 0.233) versus 73.3% to 86.7% for LLMs (kappa = 0.600 to 0.800). Healthdirect operated under conservative interactive defaults while LLMs received complete information in a single prompt, which may have disadvantaged Healthdirect. Emergency sensitivity was 46.7% versus 80.0% to 86.7% for LLMs. Healthdirect produced two critical misses; no LLM produced any across 270 evaluations (95% CI 0% to 1.4%). When LLMs undertriaged, they recommended GP care rather than self-care. No tier differences were significant (all p > 0.05), and most systems over-triaged self-care cases. Interpretation: Frontier LLMs demonstrated higher triage accuracy and safer error profiles than Healthdirect. All LLMs avoided critical misses; Healthdirect did not. Premium subscriptions did not significantly improve triage safety. These findings support clinical governance decisions about whether LLMs warrant formal evaluation alongside government-backed symptom checkers.
Amancio, R. T.; Cruz, L. N.; Dantas, R. d. S.; Gomes, M. P.; Silva, A. d. A. B. d.; Brasil, P. E.
Show abstract
Background: Health and administrative professionals in tertiary hospitals face high levels of occupational stress, mental illness, and multimorbidity. The integrated measurement of these multidimensional health aspects is essential for informing effective workplace health promotion strategies. Objective: To describe the general health status of federal public hospital staff and correlate the measured health dimensions to inform institutional health promotion initiatives. Methods: A cross-sectional, online survey study was conducted at Hospital Universitario dos Servidores do Estado (HUSE) between November and December 2025. Data collection was performed via online REDCap questionnaires covering sociodemographic profiles and validated instruments (SRQ-20, MIDAS, AUDIT, WHOQOL-BREF, PHI, WHOQOL-SRPB BREF, CBI, GPAQ, and EPSO). Descriptive statistics, comparisons across employment ties (permanent vs. contracted staff), and Spearman correlation matrices were calculated. Results: Among 197 accesses, 117 completed the informed consent, and 86 finished all questionnaires. Participants were predominantly female, aged 40 to 60 years, and Christian. Screening positivity was 25% for common mental disorders, 21% for headache-related disability, and 10% for harmful drinking. Burnout scores clustered in the second quartile, while quality of life, happiness, and spirituality scores were in the upper third. Median physical activity was 670 min/week. Mental symptoms (SRQ-20), headache (MIDAS), and burnout (CBI) correlated positively with each other and negatively with quality of life, happiness, spirituality, and institutional support (EPSO). Conclusion: The set of instruments proved feasible for situational health diagnosis among hospital staff. Although the sample size was limited in this baseline wave, the initiative fostered workplace health awareness, driving concrete initiatives, including an on-site functional gym and workplace vaccination campaigns.
Landray, I.; Carpenter, J.; Free, C.
Show abstract
Background Preventing sexually transmitted re-infections brings health benefits and can be significantly less costly than treating their sequelae. Safetxt is a potential novel digital intervention developed to promote safer sexual behaviours. However, a recent randomised controlled trial of safetxt found no effect on reinfection at 1 year (OR 1.13, 95%CI: 0.98-1.31). We investigated if safetxt's effect was mediated through sexually risky behaviours. Methods We used data from 6248 young people with STIs from 92 UK sexual health clinics. The direct and indirect effects of safetxt on reinfection were estimated using the counterfactual approach. Condom use at last sexual encounter, number of sexual partners and STI testing were assessed as mediators. These were analysed singly and together, using regression models and a formal weighting approach. The assumptions of each approach were considered and tested. Analyses were repeated in the subgroup showing the most promising effect of safetxt: men who have sex with men or with men and women (MSM/MSMW). Results No evidence was found for the total, indirect or direct effects differing from the null. Despite not being significant, for MSM/MSMW, some of safetxt's effect on reducing reinfection was identified as being offset through its effect on number of sexual partners. Conclusions There was no evidence that safetxt's effect on reinfection was mediated through changes in sexually risky behaviours. Adaptations to specifically target these behaviours are unlikely to improve safetxt's overall effect. However, improving safetxt's effect on the number of sexual partners a participant has may improve its effect for MSM/MSMW.
Hickman, R.; Joyce, D. W.; Gray, N.; Hampshire, A.; Hellyer, P. J.; Cai, Z.; Shergill, S.; D'Oliveira, T. C.
Show abstract
Background Sleep, mood, and affective states are mutually connected. There is a paucity of studies, however, that have considered bidirectional relationships between daily sleep-affective dyads in naturalistic settings, particularly for shift workers. Objective To evaluate the dynamic and temporal interplay of daily smartphone-based self-reported sleep measurements, dimensions of affective experience and cognitive processing in UK shift working nurses. Methods The EClocker Study prospectively monitored 102 National Health Service (NHS) nurses (aged 25-61 years, 83.3% female) working standard (day shift) and non-standard (fast rotating shifts) schedules over a two-week period. Smartphone-based Experience Sampling Methodology (ESM) recorded daily sleep, mood, momentary affect and cognitive attentional functioning. Self-reported burnout, emotional dysregulation, emotion reactivity and affective dimensions (positive and negative) were also collected. Findings Overall, NHS nurses reported a high prevalence of depressive symptoms, stress, burnout and sleep-circadian rhythm disturbances. Generalised Additive Modelling (GAMs) revealed that NHS nurses higher perceived sleep quality predicted better next-day mood state, while better daytime mood was associated with reduced sleep onset latency, such that participants reported falling asleep faster. In contrast, daytime mood or affect (positive and negative) had no substantial, direct impact on nurses subjective sleep parameters (sleep quality, sleep duration, sleep efficiency). Exposure to fast rotating night shifts across the two-week study was associated with more frequent response errors on a Choice Reaction Time (CRT) cognitive task, while daytime somnolence did not adversely influence nurses momentary reaction time speeds or attentional function. Conclusions Clinically relevant sleep impairments, insomnia-related symptoms, elevated stress, and poor mood were pervasive in a sample of UK NHS nurses, regardless of shift type. Sleep quality impacted next-day mood and daytime mood impacted sleep latency, while rotating shifts led to an increase in cognitive errors. Recognising the impact of shiftwork and designing interventions to promote better sleep quality offer potential to enhance mood and performance in healthcare professionals. Clinical implications We need to implement and evaluate interventions that regularise sleep patterns and promote sleep quality to alleviate mood symptoms among frontline NHS shift workers.
Laigaard, J.; Moeller, M. O.; Olsen, M. H.; Overgaard, S.; Mathiesen, O.; Karlsen, A. P. H.
Show abstract
Background: In Denmark, perioperative high-dose glucocorticoid treatment were step-wisely implemented for total hip arthroplasty (THA), total knee arthroplasty (TKA), and unicompartmental knee arthroplasty (UKA). We aimed to estimate the effect of a single high dose of glucocorticoids on opioid consumption following primary THA, TKA, and UKA. Methods: This was a prespecified analysis of a multicenter natural experiment using electronic health record data. We included elective THA, TKA, or UKA surgeries performed in Eastern Denmark from 2017-2025. At each center, surgeries before implementation of high-dose glucocorticoids served as controls, whereas surgeries after implementation comprised the intervention group. The primary outcome was the between-group difference in cumulative 0-24h opioid consumption, which included preemptive end-of-surgery doses. The predefined minimal important difference was set at 5 mg IV morphine equivalents. Secondary outcomes were maximum 0-10 numerical rating scale (NRS) pain score and incidence of opioid-related adverse events within 24 hours, hospital length of stay, and days alive and out of hospital at 30 days. Results: A total of 47,317 surgeries performed at nine centers were analyzed: 13,010 controls and 34,307 in the intervention group. During the study period, five centers implemented high-dose glucocorticoids for THA patients, two for TKA/UKA patients. High-dose glucocorticoids were administered to 6% of patients before implementation versus 92% after. High-dose glucocorticoids resulted in a mean reduction of 3.8 mg intravenous (IV) morphine equivalents (95% CI 3.3;4.3). The intervention also reduced the maximum 0-24h NRS pain score by 0.8 points (99% CI 0.7;0.9), but there was no difference in adverse events, length of stay, or days alive and out of hospital. Conclusions: Implementation of high-dose glucocorticoids reduced 0-24-hour opioid consumption by 3.8 mg IV morphine equivalents after elective hip and knee arthroplasty. This difference was below the prespecified minimal important difference threshold. Online registration: https://doi.org/10.1101/2025.11.11.25339982
Yao, R.; Wi, C.-I.; Beenken, M. J.; Watson, D.; Wheeler, P. H.; Finch, M.; Kelleher, D. P.; Anil, G.; Anderson, T.; Madden, K.; Okuno, S. H.; Odedina, F. T.; Westfall, E. C.; Park, E. Y.; Sharma, P.; Dugani, S.; Foss, R. M.; Hidaka, B. H.; Sosso, J. L.; Sabarish, S.; Singh, G.; Lugo-Fagundo, N.; Howick, J.; Kim, W. R.; Calvin, A. D.; Walker-Mcgill, C. L.; Rennert, L.; Juhn, Y. J.; Cerhan, J. R.; Lynch, B. A.
Show abstract
Purpose: This study assesses the association between colorectal cancer (CRC) screening and a validated, housing-based measure of individual-level socioeconomic status (SES, called HOUSES hereafter) within rural communities and determines whether HOUSES-integrated geospatial analysis can be used to tailor interventions. Methods: We used CRC screening data from a subset of Mayo Clinic Midwest patients living in cities without ready access to routine care in the Mayo Clinic Health System in 2019 to represent rural communities. At the individual level, we assessed the association between CRC screening rates and the HOUSES index, adjusting for age, sex, race/ethnicity, comorbidity, distance from home address to clinic, and area deprivation index, using a multilevel mixed-effects logistic regression model. Additionally, we conducted geospatial analysis to examine the correlation between hotspots of 1) lower CRC screening rates and 2) lower SES of the subject population (HOUSES quartile 1). Findings: Among 34,489 individuals (median age 64.0 years, 52.4% female), those with the lowest SES (HOUSES Q1) had 37% lower odds of being CRC screening adherent than those with the highest SES (HOUSES Q4) (adj. OR [95% CI]: 0.63 [0.58-0.69]). In the 14 identified HOUSES Q1 hotspots, there was a significant correlation in counts of HOUSES Q1 and low CRC screening (correlation coefficient=0.81). Conclusion: Lower SES was significantly associated with lower CRC screening among rural populations. HOUSES-enabled geospatial analysis identified geographic hotspots with lower CRC screening rates for targeted interventions to address disparities in CRC screening in rural communities. HOUSES may be a useful digital tool for cancer preventive care and research.
Rabbani, N.; Mettner, J.; Lee, K.; Soto-Rivera, C. L.; Windberger, A.; Santiago, K.; Hatoun, J.; Correa, E. T.; Vernacchio, L.; Kohane, I.
Show abstract
Routine childhood growth surveillance is a cornerstone of pediatric care. Growth pattern abnormalities are often early manifestations of chronic disease. Yet subtle abnormalities are frequently underrecognized, leading to diagnostic delays and avoidable morbidity. We introduce SPROUT (System for Pediatric Recognition Of Undiagnosed Trajectories), a generalized, multi-agent large language model (LLM) reasoning system designed to identify a broad spectrum of pediatric growth-related conditions from longitudinal electronic health records (EHRs) earlier than standard clinical practice. Using a large pediatric primary care EHR dataset, we developed and validated SPROUT as a two-stage system. First, a highly specific LLM screener flags concerning longitudinal growth patterns. Second, an Orchestrator module coordinates a multidisciplinary panel of LLM agents to generate a ranked differential diagnosis. To correct systemic reasoning errors, a Trainer module injects meta-knowledge into the panel via a dedicated "Learner" agent. Diagnostic capability was evaluated using a walk-forward, visit-by-visit simulation leading up to the diagnosis date. The SPROUT screener model achieved 98% (83/85) specificity and 28% (9/32) sensitivity on a gold-standard dataset of pediatric primary care patients when evaluated one year before the index date, and 100% specificity and 47% sensitivity when evaluated using longitudinal data up to the day of diagnosis. When applied to 300 control patients (i.e., healthy or undiagnosed), the screener flagged 15. Subsequent expert panel review confirmed high suspicion for undiagnosed pathology in 33% (5/15) of these cases. In chronological walk-forward validation on disease cases, the diagnostic engine identified conditions well before standard-of-care documentation. One year prior to clinical diagnosis, the system achieved sensitivities of 81% for type 1 diabetes mellitus, 56% for pituitary disorders, and 44% for celiac disease. The SPROUT multi-agent system demonstrates the ability to detect a significant portion of latent growth-related pediatric conditions months to years before current clinical standards while minimizing false positives. These results support its potential as a decision support tool for reducing diagnostic delays in pediatric care.
Song, Q.; Ni, C.; Liu, W.; Li, Y.; Malin, B. A.; Yin, Z.
Show abstract
Automatic coding from clinical notes has been studied extensively for International Classification of Diseases (ICD) codes, yet broad Current Procedural Terminology (CPT) and Healthcare Common Procedure Coding System (HCPCS) recommendation remains comparatively underexplored. Existing studies often focus on one specialty, a limited code vocabulary, or a single model family, leaving it unclear how different artificial intelligence (AI) paradigms perform under a common, clinically meaningful evaluation. We formulate CPT and HCPCS coding as an AI-assisted recommendation task in which a physician or professional coder reviews a short, ranked list of candidate codes supported by the clinical note. Using operative notes from Vanderbilt University Medical Center (VUMC) and discharge summaries from Medical Information Mart for Intensive Care IV (MIMIC-IV), we compare lexical retrieval, Clinical-Longformer, GPT-5.6-Sol, MedGemma-27B, and an inspectable agentic-style retrieve-and-verify system under a controlled review budget. Micro-averaged recall within a fixed number of recommendations measures whether reference codes reach the reviewable list; micro-F1 is reported only where reference labels are sufficiently complete. Zero-shot GPT-5.6-Sol achieves the highest recall within five and ten candidates: 0.717 and 0.800 on VUMC and lower-bound values of 0.689 and 0.738 on MIMIC-IV. The retrieve-and-verify system reaches 0.695 and 0.784 on VUMC and lower-bound values of 0.575 and 0.657 on MIMIC-IV, with a candidate-linked evidence window attached to each retained recommendation. Diagnostic analyses reveal distinct failure sources, including output-length underfilling, confusion among closely related codes, out-of-knowledge-base generation, and incomplete evidence support. These findings establish a systematic evaluation framework for procedure-code recommendation and identify practical requirements for future systems that are accurate, review-efficient, and grounded in clinical evidence.
Davis, J. T.; Kaur, G.; Hines, A.; Ben-Nun, M.; Venkatramanan, S.; Brooks, L.; Mathis, S.; Ajelli, M.; Litvinova, M.; Kummer, A. G.; Ventura, P. C.; Mhade, S.; Weber, D.; Shemetov, D.; DeFries, N.; McDonald, D. J.; Yamana, T.; Zepeda-Tello, R.; Shaman, J.; Yaari, R.; Pei, S.; Webber, A.; Shandross, L.; Ray, E.; Wadsworth, S.; Niemi, J.; Redman, W. T.; Mullany, L.; Posner, R.; Mallela, A.; Lin, Y. T.; Hlavacek, W. S.; Smart, A.; Gill, A. A.; Drennan, A.; Fiebiger, B. J.; Miller, E. F.; Lee, J.; Mihaljevic, J. R.; Geist, K. A.; Baltz, M.; Bernik, O.; Truong, Y.-M. B.; Chen, Y.; Grosvenor, C. J.;
Show abstract
Forecasting influenza hospitalizations informs public health preparedness, yet questions remain about which types of forecasts best guide action. We evaluate categorical trend forecasts, which communicate probabilities of upcoming increases or decreases in epidemic trajectories, submitted to CDC's FluSight Forecasting Challenge between Fall-2024 and Spring-2026. Teams submitted probability distributions over five categories describing direction and magnitude of week-over-week changes in laboratory-confirmed influenza hospital admissions. We assessed performance using Ranked Probability Skill Score, Brier Skill Score, and measures of forecast-observation agreement. Most models outperformed an equal-probability baseline; the FluSight ensemble ranked among the top three in the 2024-25 and 2025-26 seasons. Forecasts were most accurate during stable periods and least during periods of rapid change, with most models underestimating observed trends. Conclusions were robust to choice of scoring metric and reference model. These results support categorical trend ensembles as an approach to communicating infectious disease forecasts that may inform public health decision-making.
de Araujo Morais, J. H.; Dias Ferreira, C.; Saraceni, V.; Medeiros de Oliveira Cruz, D.; Mateus Oliveira Aguilar, G.; Cruz, O. G.
Show abstract
Motivation: With the scaling frequency and intensity of extreme heat events across the globe, it is critical for public institutions to develop early detection systems and continuous monitoring of these events and their impacts. In Brazil, Rio de Janeiro was the first city to publish its heat protocol, with the Rio Heat Dashboard as a central component of this system. Implementation: The dashboard was implemented using R/Shiny and integrates climatic and health data from multiple sources. General features: The application comprises real-time heat exposure monitoring and automatic alert level classification, which is monitored daily by multiple municipal actors and supports activation of actions specified in the heat protocol. It also features a health impact module, which lists each heat event and its impact on mortality, and primary care and emergency visits. Availability: The source for full reproducibility is available through https://github.com/joaohmorais/RioHeatDashboard.
Bandini, V.; Whitaker, L. H.; Vincent, K.; Salmeri, N.; Mawson, R.; Vercellini, P.; Horne, A. W.
Show abstract
Background: Endometriosis is a chronic pain condition in which hormonal therapies form the cornerstone of long-term management. Treatment tolerability is critical for adherence and therapeutic success, but most comparative studies and reviews have focused on their ability to reduce menstrual pain, while their impact on non-menstrual pelvic pain (NMPP), bleeding patterns, adverse events (AEs), treatment discontinuation and quality of life (QoL) remain poorly characterised. This systematic review and meta-analysis evaluate these outcomes across currently available hormonal therapies, providing practical evidence for clinical decision-making. Methods: PubMed/MEDLINE, Scopus, and Embase were searched up to November 2025 for randomised controlled trials comparing at least two active first- or second-line hormonal treatments for endometriosis. Studies without confirmed endometriosis, treatment duration less than three months and comparing therapies to placebo only were excluded. Data were extracted by two reviewers from reports. Pain outcomes were pooled as mean differences (MD, 95% CI), with bleeding patterns, AEs, and discontinuations as proportions. Analyses were performed in R. PROSPERO: CRD420251137785. Findings: Of 1892 records screened, 48 trials (5583 women) met our inclusion criteria. Overall pelvic pain (0-10 scale) was significantly reduced across all treatment categories (p<0.001): combined oral contraceptives (COCs) (MD 3.17), oral and long-acting progestogens (MD 3.83; MD 4.29), and GnRH-analogues (MD 3.81). Sensitivity analyses restricted to studies reporting NMPP yielded comparable results. GnRH-agonists showed the most favourable bleeding profile, followed by continuous COCs. However, all regimens reported class-specific AEs, including mood changes, nausea, headache, weight gain, and decreased libido (pooled proportions >10%). Overall discontinuation due to AEs was 7.7%, and vaginal bleeding was the leading cause. Heterogeneity across meta-analyses was high. Risk of bias (RoB2) was moderate to high. Interpretation: Given similar reductions in overall pelvic pain across hormonal therapies, treatment decisions should prioritise differences in bleeding profiles, therapy-specific AEs, and QoL. Funding: None.
Knol, L.; Nagpal, A.; Hussain, F.; Beckmann, C. F.; Leow, A.; Eisenlohr-Moul, T. A.; Marquand, A. F.
Show abstract
Digital phenotyping, which is defined as quantifying someone's behaviour with digital devices, provides unprecedented opportunities for understanding human mental health but is hampered by high levels of inter-individual variability. Here, we propose a new method to address this, parsing inter-individual variability by decomposing the digital phenotype dynamics into latent trajectories and using each individual's trajectory membership as a moderator when modelling psychopathology over the same timeframe. We applied our method in the context of mood symptom exacerbation across the menstrual cycle, where symptom severity and timing are inconsistent between individuals. Using the BiAffect platform to collect smartphone typing dynamics, we found stable trajectories in smartphone movement rate: one group of participants showed substantial movement rate fluctuations across the menstrual cycle, whilst the others did not. Participants with movement fluctuations displayed increased fluctuations across the cycle in prospective anhedonia and depression ratings, but not in anxiety, irritability, and suicidal ideation.